Client Registry / Master Patient Index
The client registry answers one question: are these two records about the same person? Everything else in a health information exchange depends on the answer being right.
The terms differ slightly by tradition — MPI (master patient index) is the hospital-sector term, client registry the OpenHIE term, EMPI (enterprise MPI) the term when it spans organisations — but the machinery is the same.
What it holds
A small, deliberately constrained record per person:
- Identifiers, each with its issuing namespace (national ID, health ID, facility MRNs)
- Name, in the forms the culture actually uses
- Date of birth, with a flag for estimated dates
- Sex, and separately gender where relevant
- Contact — phone, address, and their history
- Mother's name, or another culturally appropriate discriminator
- Links to source system records, and the confidence of each link
- Merge history
It holds no clinical data. A client registry that starts storing diagnoses has become a shared health record and inherits all of that component's governance requirements.
Matching
Deterministic matching
Exact agreement on a defined set of fields, usually as an ordered rule set:
Rule 1: national_id matches exactly → MATCH
Rule 2: health_id matches exactly → MATCH
Rule 3: surname + given name + date_of_birth + sex all match → MATCH
Otherwise → NO MATCH
Fast, explainable, auditable. A clerk can be told why two records matched, and a court can be shown the rule.
Its weakness is that it fails on exactly the data that health systems have: transliterated names, estimated birth dates, single-name populations, phonetic spelling variation, and identifiers that are absent or wrong.
Probabilistic matching
Each field contributes a weight based on how much agreement or disagreement on that field shifts the odds. Weights are summed into a score.
The insight from Fellegi–Sunter record linkage theory: a field's evidential value depends on how discriminating it is. Agreeing on an uncommon surname is strong evidence; agreeing on a common one is weak. Agreeing on sex is worth almost nothing, because half the population agrees by chance.
score
│
│ no match review match
├──────────────┬──────────────┬───────────────────▶
0 lower upper
threshold threshold
- Above the upper threshold: automatic link
- Below the lower threshold: automatic non-link
- Between: human review
Techniques: string similarity (Jaro-Winkler, Levenshtein), phonetic encoding (Soundex, Double Metaphone, and language-appropriate equivalents — note that Soundex is built for English and performs badly on many languages), date tolerance, nickname and transliteration dictionaries, blocking to avoid comparing every record with every other.
Choosing
| Deterministic | Probabilistic | |
|---|---|---|
| Explainability | High | Moderate — needs score breakdowns |
| Performance on clean data | Excellent | Excellent |
| Performance on messy data | Poor | Better |
| Tuning effort | Low | Ongoing |
| Governance burden | Low | Requires a review team and threshold policy |
Start deterministic if data quality does not support anything else, and record that as the reason. Probabilistic matching over data with 40% missing birth dates does not produce better matches; it produces confident wrong ones. Revisit the decision when the data improves — that is what the ADR is for.
Most mature registries are hybrid: deterministic rules on strong identifiers, probabilistic scoring for the rest.
Thresholds are a clinical safety decision
Two error types, with asymmetric consequences:
| Error | Consequence |
|---|---|
| False positive — two people merged | One person's allergies, results and medications appear in another's record. Direct patient-harm risk. |
| False negative — one person split | Fragmented record; duplicate testing; missed history; inflated patient counts. |
False positives are worse. Set thresholds conservatively, route uncertainty to human review, and make the review queue a funded role rather than a background task. The single best predictor of whether an MPI programme succeeds is whether anyone actually works the review queue.
Measure and publish: auto-match rate, review queue depth and age, duplicate rate in source systems, and — periodically, against a manually adjudicated sample — false positive and false negative rates.
The golden record
The registry's consolidated view of a person, assembled from source records.
Survivorship rules decide which value wins when sources disagree: most recent, most trusted source, most complete, or human-adjudicated. Record the rule per field, and retain the source values — the golden record is a view, not a replacement. When a merge turns out to be wrong, the source values are what make recovery possible.
Merge and unmerge
Record A ──┐
├──▶ Golden record (survivor) A and B remain, marked as
Record B ──┘ replaced-by, with full history
Requirements:
- The losing identifier must continue to resolve — old references in old
systems will keep arriving for years. FHIR handles this with
Patient.linkand thereplaced-bytype. - Downstream systems must be notified, and must be able to act on it. This is the hardest part: the shared health record, the HMIS, the insurer and every EMR each hold data keyed to the old identifier.
- Unmerge must be possible. Merges are sometimes wrong, and discovering that after six months of clinical data has accumulated under the merged identity is a genuine emergency. If your platform cannot unmerge, that is a decisive selection criterion.
- Every merge and unmerge is audited with the evidence and the deciding person.
The FHIR interface
POST /Patient/$match # find candidates for a demographic set
GET /Patient?identifier=… # look up by identifier
GET /Patient/123/$everything # (on the SHR, not the registry)
$match is the standard operation and returns candidates with a match grade
(certain, probable, possible, certainly-not) and a score. IHE PIX and
PDQ profiles provide the equivalent in the pre-FHIR world and are still
widespread.
Where it sits
Registration desk ─────┐
CHW app ───────────────┼──▶ Interoperability layer ──▶ Client registry
Laboratory ────────────┘ │
┌──────────┴──────────┐
▼ ▼
Golden record Review queue
│ (humans)
▼
shared identifier returned to callers
The registry resolves identity; it does not store the encounter. See interoperability layer.
Practical guidance
- Fix data capture before tuning the algorithm. A registration form that accepts a free-text birth date generates more duplicates than any matcher can repair. Enforce format, use pickers, add check digits, and show the clerk likely existing matches before they create a new record.
- Prevent duplicates at the point of creation. A search-before-create step is worth more than any downstream reconciliation.
- Load real data early. Matching performance is a property of your population's naming and data-quality patterns, not of the software.
- Test with adversarial cases: twins, common names, single-name populations, transliteration variants, name changes at marriage, estimated birth dates, and family members sharing a phone number.
- Publish the duplicate rate and hold source systems accountable for it.
- Plan the review team before go-live, including who covers it.
Open-source options
| Project | Notes |
|---|---|
| OpenCR | The OpenHIE community client registry; FHIR-based, configurable deterministic and probabilistic rules |
| SanteMPI / SanteDB | Full-featured MPI and health data platform, IHE and FHIR interfaces |
| OpenEMPI | Long-established; verify current maintenance before adopting |
| Splink | Not health-specific: an open-source probabilistic linkage library, useful for evaluating matching strategies against your own data before committing |
All Tier 2. See platforms.
References
- OpenHIE client registry — https://ohie.org/
- FHIR
Patient— https://hl7.org/fhir/patient.html - FHIR
$matchoperation — https://hl7.org/fhir/patient-operation-match.html - IHE PIX/PDQ — https://www.ihe.net/
- Fellegi & Sunter, A Theory for Record Linkage (1969), JASA 64(328) — the basis for probabilistic matching
- Splink — https://moj-analytical-services.github.io/splink/